Papers with data extraction
DepressMind: A Depression Surveillance System for Social Media Analysis (2024.eacl-demo)
Copied to clipboard
| Challenge: | DepressMind is a tool for the analysis of social network data on depression . the tool explores multiple psychological dimensions associated with clinical depression based on the social network . |
| Approach: | They propose to use social network data to analyze clinical depression . they aim to link extracts from social networks with symptoms of the Beck Depression Inventory . |
| Outcome: | The tool explores multiple psychological dimensions associated with clinical depression and estimates the extent to which these symptoms manifest in language use. |
Exploiting the Shadows: Unveiling Privacy Leaks through Lower-Ranked Tokens in Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models face vulnerabilities related to the extraction of sensitive information. |
| Approach: | They propose a method to exploit the model's lower-ranked output tokens to extract private information from retrieved documents or training knowledge. |
| Outcome: | The proposed method is effective in both the agentic application privacy extraction setting and the direct training data extraction. |
LLM-Based Web Data Collection for Research Dataset Creation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | researchers across many fields rely on web data to gain new insights and validate methods. |
| Approach: | They propose a human-in-the-loop framework that automates web-scale data collection end-to-end using large language models (LLMs) |
| Outcome: | The proposed framework outperforms existing methods in three different tasks and a user evaluation demonstrates its practical utility. |
ETHICIST: Targeted Training Data Extraction Through Loss Smoothed Soft Prompting and Calibrated Confidence Estimation (2023.acl-long)
Copied to clipboard
| Challenge: | Recent studies show that pre-trained language models memorize a considerable fraction of training data, leading to privacy risk of information leakage. |
| Approach: | They propose a method for targeted training data extraction using a smoothed soft prompting and calibrated confidence estimation. |
| Outcome: | The proposed method significantly improves the extraction performance on a recently proposed public benchmark. |
A Manually Annotated Resource for the Investigation of Nasal Grunts (2020.lrec-1)
Copied to clipboard
| Challenge: | acoustic annotation of nasal grunts is described in the whole CID corpus of the french language . acculturation of non-lexical conversational sounds has been debated for a long time . |
| Approach: | They propose an annotation framework for nasal grunts of the whole French CID corpus . they characterise acoustic cues and visual cue conventions followed for the annotation . |
| Outcome: | The proposed framework is based on the entire French CID corpus. |
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Recent surveys of literature highlight the overwhelming growth of Large Language Models (LLMs). |
| Approach: | They propose a semi-automated literature analysis approach that automates literature analysis using LLMs. |
| Outcome: | The proposed approach reduces paper surveying and data extraction by 93% compared to manual methods. |